Audio-visual speech synthesis for finnish
نویسندگان
چکیده
We describe our Finnish audio-visual speech synthesizer, its evaluation and discuss possible improvements. We have combined a three dimensional facial model with a commercial audio text-to-speech synthesizer. The visual speech is based on a letter-to-viseme mapping and the animation is created by linear interpolation between the visemes. An intelligibility test was run to quantify the benefit of seeing the synthetic and natural face on hearing the synthetic and natural voice presented at different signal to noise ratios. Both natural and synthetic faces improved the intelligibility of both natural and synthetic auditory speech. We examined the confusion patterns of consonants and the identification of the Finnish visemes. We also propose how the viseme repertoire of the talking head can be improved.
منابع مشابه
INTERSPEECH 2006 1 sing Dominance Functions and udio - Visual Speech Synthesis
This paper presents results of training of coarticulation models for Czech audio-visual speech synthesis. Two approaches for solution of coarticulation in audio-visual speech synthesis were used, coarticulation based on dominance functions and visual unit selection. For both approaches, coarticulation models were trained. Models for unit selection approach were trained by visualy clustered data...
متن کاملText-to-audio-visual speech synthesis based on parameter generation from HMM
This paper describes a technique for synthesizing auditory speech and lip motion from an arbitrary given text. The technique is an extension of the visual speech synthesis technique based on an algorithm for parameter generation from HMM with dynamic features. Audio and visual features of each speech unit are modeled by a single HMM. Since both audio and visual parameters are generated simultan...
متن کاملMediaTeam Speech Corpus: a first large Finnish emotional speech database
In this paper, a large Finnish emotional speech database is introduced. The database contains simulated emotional speech, reflecting some of the affective states commonly known as basic emotions. The database has been analyzed instrumentally in terms of some 40 acoustic/prosodic parameters, and is currently being annotated in terms of a number of linguistically relevant contrasts. The annotatio...
متن کاملInnovations in Czech audio-visual speech synthesis for precise articulation
This paper presents new steps toward animation of precise articulation. The acquisition of audio-visual corpus for Czech and new method for parameterization of visual speech was designed to obtain exact speech data. The parameterization method is primarily suitable for training a data driven visual speech synthesis systems. The audio-visual corpus includes also specially designed test part. Fur...
متن کاملCzech Audio-Visual Speech Synthesis with an HMM-trained Speech Database and Enhanced Coarticulation
The task of visual speech synthesis is usually solved by concatenation of basic speech units selected from a visual speech database. Acoustical part is carried out separately using similar method. There are two main problems in this process. The first problem is a design of a database, that means estimation of the database parameters for all basic speech units. Second problem is a way how to co...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 1999